Welcome to Implementing High-Availability PostgreSQL Clusters via Patroni. Database downtime means application downtime. Standard PostgreSQL replication provides read replicas but lacks native, robust automated failover. Patroni has emerged as the industry standard template for building HA PostgreSQL solutions that failover automatically and prevent split-brain scenarios.
1. The Anatomy of Patroni
Patroni is an open-source Python daemon co-developed by Zalando that runs alongside PostgreSQL. It relies on a Distributed Configuration Store (DCS)βusually etcd, Consul, or ZooKeeperβto maintain the cluster's state. By pushing the state into a consensus algorithm (Raft), Patroni ensures that all nodes agree on who the primary database is.
2. How Automated Failover Works
In a typical 3-node setup (one Primary, two Replicas), the Primary holds a "leader lock" in etcd with a TTL (Time To Live). The Patroni daemon on the Primary continually renews this lock. If the Primary server crashes or loses network connectivity, the lock expires.
The remaining Replicas detect the missing leader lock. They then compare their WAL (Write-Ahead Log) positions. The Replica with the most up-to-date transaction log acquires the lock, promotes itself to Primary, and the other node automatically reconfigures itself to replicate from the new Primary. This entire process typically completes in under 30 seconds.
3. Handling Split-Brain and Fencing
A "split-brain" occurs when a network partition makes two nodes think they are the Primary, resulting in corrupted data. Patroni prevents this inherently through the DCS. Because etcd requires a strict majority (quorum) to write data, a node isolated from the network cannot acquire the leader lock.
However, an isolated node might still believe it is the primary for a few seconds. Patroni supports "fencing" mechanisms (like terminating the Postgres process or interacting with IPMI to power off the node) to guarantee the old primary cannot accept writes.
4. Routing Traffic with HAProxy
Applications need to know where to send read and write queries. Patroni provides a REST API that returns a 200 OK if the node is the Primary, and a 503 Service Unavailable if it's a Replica. By configuring HAProxy with httpchk against this API endpoint, you can automatically route all write traffic (Port 5432) exclusively to the current Primary, instantly updating if a failover occurs.
5. Disaster Recovery Integration with pgBackRest
While HA protects against hardware failure, it does not protect against accidental data deletion (e.g., DROP TABLE). Patroni integrates tightly with pgBackRest for Point-in-Time Recovery (PITR). Replicas can be configured to fetch WAL files from an S3 bucket via pgBackRest if they fall too far behind the primary to sync via streaming replication.
Conclusion
Building a Patroni cluster requires orchestrating PostgreSQL, Python, HAProxy, and etcd. While complex, it creates a self-healing database architecture capable of surviving catastrophic node failures with zero human intervention.